Goto

Collaborating Authors

 spcas9 activity


A deep learning-based model DeepSpCas9 to predict SpCas9 activity

#artificialintelligence

In a new report on Science Advances, Hui Kwon Kim and interdisciplinary researchers at the departments of Pharmacology, Electrical and Computer Engineering, Medical Sciences, Nanomedicine and Bioinformatics in the Republic of Korea, evaluated the activities of SpCas9; a bacterial RNA-guided Cas9 endonuclease variant (a bacterial enzyme that cuts DNA for genome editing) from Streptococcus pyogenes. They used a high-throughput approach with 12,832 target sequences based on a human cell library to build a deep learning model and predict the activity of SpCas9. The data contained oligonucleotides (nucleotides or building blocks) containing target sequence pairs and a corresponding guide sequence to encode single-guide RNA (sgRNA), which can direct the Cas9 protein to bind and cleave a specific DNA sequence for genome editing. They implemented deep learning-based training on the large dataset of SpCas9-induced indel (insertion or deletion) frequencies to develop an SpCas9 activity predicting model named DeepSpCas9 now available online. When the team tested the software against independently generated datasets, the results showed high generalization performance, i.e. the model could properly adapt to new, previously unseen data.


SpCas9 activity prediction by DeepSpCas9, a deep learning–based model with high generalization performance

#artificialintelligence

To increase the accuracy of the analysis, deep sequencing data were filtered; target sequences with deep sequencing read counts below 200 and background indel frequencies above 8% were excluded as similarly performed previously (21). DNase-sequencing (DNase-seq) narrow peak data from ENCODE (36) were used to calculate chromatin accessibility as previously described (21). For each target site, 23 bases of the PAM plus protospacer sequence were aligned to the hg19 human reference genome using bowtie (41). Only the target sites that overlapped with DNase-seq narrow peaks were considered as DNase I hypersensitive target sites. We divided the Endo_Cas9 dataset into paired subsets by stratified random sampling from strata of DHS and non-DHS sites so that a similar ratio of DHS/non-DHS sites was assigned to each subset.